Papers with large-scale training data

5 papers
Machine Comprehension Improves Domain-Specific Japanese Predicate-Argument Structure Analysis (D19-58)

Copied to clipboard

Challenge: a lack of gold datasets and knowledge about PAS analysis makes it difficult to create accurate PAS analyses.
Approach: They construct a Japanese blog-QA dataset and a reading comprehension QA dataset using crowdsourcing.
Outcome: The proposed method is most effective, pre-training model to acquire domain knowledge and fine-tuning model based on PAS-QA dataset.
InfiMM: Advancing Multimodal Understanding with an Open-Sourced Visual Language Model (2024.findings-acl)

Copied to clipboard

Challenge: InfiMM is a multimodal large language model that adapts to complex vision-language tasks.
Approach: They present a Multimodal Large Language Model that adapts to intricate vision-language tasks using large-scale training data and comprehensive training strategies.
Outcome: Empirical evaluations across a variety of benchmarks underscore InfiMM’s remarkable capability in multimodal understanding.
Language-to-Space Programming for Training-Free 3D Visual Grounding (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for 3D visual grounding have been proposed, but they are limited by the scarcity of 3D vision-language datasets and the high cost of annotations.
Approach: They propose a method for training-free 3D visual grounding that uses LLM-generated codes to analyze 3D spatial relations among objects.
Outcome: The proposed method achieves 52.9% accuracy on the Nr3D benchmark and significantly reduces grounding time and token costs.
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction (2026.findings-acl)

Copied to clipboard

Challenge: Existing MLLMs are computationally expensive and may produce hallucinated content . RA-RRG uses large language models to generate radiology reports .
Approach: They propose a retrieval-augmented RRG framework that combines multimodal retrieval with large language models to generate radiology reports.
Outcome: RA-RRG uses large language models to generate radiology reports . it suppresses hallucinations while maintaining strong report generation performance .
Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for detecting LLM-generated text require no training data.
Approach: They propose a black-box zero-shot detection approach that calculates the Grammar Error Correction Score for a given text to differentiate between human-written and LLM-generated texts.
Outcome: The proposed method outperforms current state-of-the-art zero-shot and supervised methods, achieving an average AUROC of 98.62% across XSum and Writing Prompts datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations